Model matrix
Add Qwen, Llama, Claude, GPT, Gemini, Mistral, and domain-specific model families under the same directory convention.
Open Longitudinal Benchmark Registry
prompt-cache-bench is designed as a long-running, reproducible registry for LLM inference systems. It will accumulate hundreds of model/provider A/B comparisons, raw datasets, report pages, figures, and runnable experiment code under a consistent methodology.
Each row represents a versioned benchmark artifact. Future experiments should be appended here with stable report URLs, raw datasets, reproducible scripts, and manifest checksums.
| Status | Model | A/B Pair | Design | Reports | Data |
|---|---|---|---|---|---|
| Published | google/gemini-3-flash-preview |
Infron vs OpenRouter | 4x50 paired streaming Chat Completions; routing sort modes; default reasoning/thinking; prompt-length tiers; /v1/chat/completions only. | EN HTML · ZH HTML · EN MD · ZH MD · Reports | Dataset |
| Published | qwen/qwen3.6-flash |
Infron vs OpenRouter | 4x50 paired streaming Chat Completions; routing sort modes; default reasoning/thinking; prompt-length tiers; /v1/chat/completions only. | EN HTML · ZH HTML · EN MD · ZH MD · Reports | Dataset |
| Published | qwen/qwen3.6-35b-a3b |
Infron vs OpenRouter | 4x50 paired streaming Chat Completions; routing sort modes; default reasoning/thinking; prompt-length tiers; /v1/chat/completions only. | EN HTML · ZH HTML · EN MD · ZH MD · Reports | Dataset |
| Published | google/gemma-4-26b-a4b |
Infron vs OpenRouter | 4x50 paired streaming Chat Completions; routing sort modes; default reasoning/thinking; prompt-length tiers; /v1/chat/completions only; OpenRouter alias google/gemma-4-26b-a4b-it. | EN HTML ZH HTML EN MD ZH MD Reports | Dataset |
| Published | minimax/minimax-m3 |
Infron vs OpenRouter | 4x50 paired streaming Chat Completions; routing sort modes; default reasoning/thinking; prompt-length tiers; /v1/chat/completions only. | EN HTML · ZH HTML · EN MD · ZH MD · Reports | Dataset |
| Published | minimax/minimax-m2.7 |
Infron vs OpenRouter | 4x50 paired streaming Chat Completions; routing sort modes; default reasoning/thinking; prompt-length tiers; /v1/chat/completions only. | EN HTML · ZH HTML · EN MD · ZH MD · Reports | Dataset |
| Published | minimax/minimax-m2.5 |
Infron vs OpenRouter | 4x50 paired streaming Chat Completions; routing sort modes; default reasoning/thinking; prompt-length tiers; /v1/chat/completions only. | EN HTML · ZH HTML · EN MD · ZH MD · Reports | Dataset |
| Published | xiaomi/mimo-v2.5 |
Infron vs OpenRouter | 4x50 paired streaming Chat Completions; routing sort modes; default reasoning/thinking; prompt-length tiers; /v1/chat/completions only. | EN HTML · ZH HTML · EN MD · ZH MD · Reports | Dataset |
| Published | moonshotai/kimi-k2.7-code |
Infron vs OpenRouter | 4x50 paired streaming Chat Completions; routing sort modes; default reasoning/thinking; prompt-length tiers; /v1/chat/completions only. | EN HTML ZH HTML EN MD ZH MD Reports | Dataset |
| Published | moonshotai/kimi-k2.6 |
Infron vs OpenRouter | 4x50 paired streaming Chat Completions; routing sort modes; default reasoning/thinking; prompt-length tiers; /v1/chat/completions only. | EN HTML ZH HTML EN MD ZH MD Reports | Dataset |
| Published | moonshotai/kimi-k2.5 |
Infron vs OpenRouter | 4x50 paired streaming Chat Completions; routing sort modes; default reasoning/thinking; prompt-length tiers; /v1/chat/completions only. | EN HTML · ZH HTML · EN MD · ZH MD · Reports | Dataset |
| Published | z-ai/glm-4.7 |
Infron vs OpenRouter | 4x50 paired streaming Chat Completions; routing sort modes; default reasoning/thinking; prompt-length tiers; /v1/chat/completions only. | EN HTML · ZH HTML · EN MD · ZH MD · Reports | Dataset |
| Published | z-ai/glm-5.1 |
Infron vs OpenRouter | 4x50 paired streaming Chat Completions; routing sort modes; default reasoning/thinking; prompt-length tiers; /v1/chat/completions only. | EN HTML · ZH HTML · EN MD · ZH MD · Reports | Dataset |
| Published | z-ai/glm-5 |
Infron vs OpenRouter | 4x50 paired streaming Chat Completions; routing sort modes; default reasoning/thinking; prompt-length tiers; /v1/chat/completions only. | EN HTML · ZH HTML · EN MD · ZH MD · Reports | Dataset |
| Published | qwen/qwen3.5-27b |
Infron vs OpenRouter | 4x50 paired streaming Chat Completions; routing sort modes; default reasoning/thinking; prompt-length tiers; /v1/chat/completions only. | EN HTML ZH HTML EN MD ZH MD Reports | Dataset |
| Published | openai/gpt-5.4-nano |
Infron vs OpenRouter | 4x50 paired streaming Chat Completions; routing sort modes; default reasoning/thinking; prompt-length tiers; /v1/chat/completions only. | EN HTML · ZH HTML · EN MD · ZH MD · Reports | Dataset |
| Published | openai/gpt-4o-mini |
Infron vs OpenRouter | 4 x 50, streaming, routing sort: throughput / price / latency / ttft, platform-default reasoning, short / medium / long prompt tiers, Chat Completions only | EN HTML · ZH HTML · EN MD · ZH MD · Reports | Dataset |
| Published | qwen/qwen3-next-80b-a3b-instruct |
Infron vs OpenRouter | 4 x 50, streaming, routing sort: throughput / price / latency / ttft, platform-default reasoning, short / medium / long prompt tiers, Chat Completions only | EN HTML · ZH HTML · EN MD · ZH MD · Reports | Dataset |
| Published | qwen/qwen3.5-plus |
Infron vs OpenRouter | 4 x 50, streaming, routing sort: throughput / price / latency / ttft, platform-default reasoning, short / medium / long prompt tiers, Chat Completions only | EN HTML · ZH HTML · EN MD · ZH MD · Reports | Dataset |
| Published | openai/gpt-5.4-mini |
Infron vs OpenRouter | 4 x 50, streaming, routing sort: throughput / price / latency / ttft, platform-default reasoning, short / medium / long prompt tiers, Chat Completions only | EN HTML · ZH HTML · EN MD · ZH MD · Reports | Dataset |
| Published | deepseek/deepseek-v4-pro |
Infron vs OpenRouter | 4 x 50, streaming, routing sort: throughput / price / latency / ttft, platform-default reasoning, short / medium / long prompt tiers, Chat Completions only | EN HTML · ZH HTML · EN MD · ZH MD · Reports | Dataset |
| Published | z-ai/glm-5.2 |
Infron vs OpenRouter | 4 x 50, streaming, routing sort: throughput / price / latency / ttft, platform-default reasoning, short / medium / long prompt tiers | EN HTML · ZH HTML · EN MD · ZH MD · Reports | Dataset |
| Published | deepseek/deepseek-v4-flash |
Infron vs OpenRouter | 4 x 50, streaming, routing sort: throughput / price / latency / ttft, platform-default reasoning, short / medium / long prompt tiers, Chat Completions only | EN HTML · ZH HTML · EN MD · ZH MD · Reports | Dataset |
| Not published | llama/*, claude/*, gpt/*, other qwen/* |
Provider matrix expansion | Matched-payload A/B, streaming TTFT, provider attribution, cost breakdown | To be published | To be published |
The homepage is organized around the registry dimensions that will scale as the benchmark grows.
Research papers from Infron related to LLM gateway evaluation, routing, and inference systems.
The project treats every benchmark as a controlled research artifact, not a one-off dashboard snapshot.
Each routing mode uses stable payload SHA256 values and sends two identical prompts per round.
Only `sort/group/round` pairs with equal response-side prompt tokens are retained.
Usage, cost, TTFT, latency, provider identifiers, and cache tokens are preserved when returned.
Reports, raw data, code snapshots, and checksums are committed together.
The registry is ready for broader model/provider coverage while preserving the same evidence standard.
Add Qwen, Llama, Claude, GPT, Gemini, Mistral, and domain-specific model families under the same directory convention.
Track cross-platform comparisons and upstream provider behavior, including routing drift over repeated runs.
Repeat experiments over time to observe cache TTL, provider fallback, price movement, latency tails, and throughput stability.
每一行代表一个带版本的 benchmark 工件。未来实验会继续追加稳定报告 URL、原始数据集、复现实验代码和 manifest checksum。
| 状态 | 模型 | A/B Pair | 实验设计 | 报告 | 数据 |
|---|---|---|---|---|---|
| 已发布 | google/gemini-3-flash-preview |
Infron vs OpenRouter | 4x50 配对 streaming Chat Completions;routing sort 模式;默认 reasoning/thinking;prompt-length tiers;仅 /v1/chat/completions。 | 中文 HTML · EN HTML · 中文 MD · EN MD · 报告目录 | 数据集 |
| 已发布 | qwen/qwen3.6-35b-a3b |
Infron vs OpenRouter | 4x50 配对 streaming Chat Completions;routing sort modes;默认 reasoning/thinking;prompt-length tiers;仅 /v1/chat/completions。 | 中文 HTML · EN HTML · 中文 MD · EN MD · 报告目录 | 数据集 |
| 已发布 | google/gemma-4-26b-a4b |
Infron vs OpenRouter | 4x50 配对 streaming Chat Completions;routing sort modes;默认 reasoning/thinking;prompt-length tiers;仅 /v1/chat/completions;OpenRouter 使用 alias google/gemma-4-26b-a4b-it。 | 中文 HTML EN HTML 中文 MD EN MD 报告目录 | 数据集 |
| 已发布 | minimax/minimax-m3 |
Infron vs OpenRouter | 4x50 配对 streaming Chat Completions;routing sort 模式;默认 reasoning/thinking;Prompt 长度分层;仅 /v1/chat/completions。 | 中文 HTML · EN HTML · 中文 MD · EN MD · 报告目录 | 数据集 |
| 已发布 | minimax/minimax-m2.7 |
Infron vs OpenRouter | 4x50 配对 streaming Chat Completions;routing sort 模式;默认 reasoning/thinking;Prompt 长度分层;仅 /v1/chat/completions。 | 中文 HTML · EN HTML · 中文 MD · EN MD · 报告目录 | 数据集 |
| 已发布 | minimax/minimax-m2.5 |
Infron vs OpenRouter | 4x50 配对 streaming Chat Completions;routing sort 模式;默认 reasoning/thinking;Prompt 长度分层;仅 /v1/chat/completions。 | 中文 HTML · EN HTML · 中文 MD · EN MD · 报告目录 | 数据集 |
| 已发布 | xiaomi/mimo-v2.5 |
Infron vs OpenRouter | 4x50 配对流式 Chat Completions;routing sort 模式;默认 reasoning/thinking;prompt 长度分层;仅 /v1/chat/completions。 | 中文 HTML · EN HTML · 中文 MD · EN MD · 报告目录 | 数据集 |
| 已发布 | moonshotai/kimi-k2.7-code |
Infron vs OpenRouter | 4x50 配对流式 Chat Completions;routing sort 模式;默认 reasoning/thinking;prompt 长度分层;仅 /v1/chat/completions。 | 中文 HTML EN HTML 中文 MD EN MD 报告目录 | 数据集 |
| 已发布 | moonshotai/kimi-k2.6 |
Infron vs OpenRouter | 4x50 配对 streaming Chat Completions;routing sort modes;默认 reasoning/thinking;prompt-length tiers;仅 /v1/chat/completions。 | 中文 HTML EN HTML 中文 MD EN MD 报告目录 | 数据集 |
| 已发布 | moonshotai/kimi-k2.5 |
Infron vs OpenRouter | 4x50 配对流式 Chat Completions;routing sort 模式;默认 reasoning/thinking;Prompt 长度分层;仅 /v1/chat/completions。 | 中文 HTML · EN HTML · 中文 MD · EN MD · 报告目录 | 数据集 |
| 已发布 | z-ai/glm-4.7 |
Infron vs OpenRouter | 4x50 配对流式 Chat Completions;routing sort 模式;默认 reasoning/thinking;Prompt 长度分层;仅 /v1/chat/completions。 | 中文 HTML · EN HTML · 中文 MD · EN MD · 报告目录 | 数据集 |
| 已发布 | z-ai/glm-5.1 |
Infron vs OpenRouter | 4x50 配对流式 Chat Completions;routing sort 模式;默认 reasoning/thinking;Prompt 长度分层;仅 /v1/chat/completions。 | 中文 HTML · EN HTML · 中文 MD · EN MD · 报告目录 | 数据集 |
| 已发布 | z-ai/glm-5 |
Infron vs OpenRouter | 4x50 配对流式 Chat Completions;routing sort 模式;默认 reasoning/thinking;Prompt 长度分层;仅 /v1/chat/completions。 | 中文 HTML · EN HTML · 中文 MD · EN MD · 报告目录 | 数据集 |
| 已发布 | openai/gpt-5.4-nano |
Infron vs OpenRouter | 4x50 配对流式 Chat Completions;routing sort 模式;默认 reasoning/thinking;Prompt 长度分层;仅 /v1/chat/completions。 | 中文 HTML · EN HTML · 中文 MD · EN MD · 报告目录 | 数据集 |
| 已发布 | openai/gpt-4o-mini |
Infron vs OpenRouter | 4 x 50,streaming,routing sort: throughput / price / latency / ttft,平台默认 reasoning,short / medium / long prompt tiers,仅 Chat Completions | 中文 HTML · EN HTML · 中文 MD · EN MD · 报告目录 | 数据集 |
| 已发布 | qwen/qwen3-next-80b-a3b-instruct |
Infron vs OpenRouter | 4 x 50,streaming,routing sort: throughput / price / latency / ttft,平台默认 reasoning,short / medium / long prompt tiers,仅 Chat Completions | 中文 HTML · EN HTML · 中文 MD · EN MD · 报告目录 | 数据集 |
| 已发布 | qwen/qwen3.5-plus |
Infron vs OpenRouter | 4 x 50,streaming,routing sort: throughput / price / latency / ttft,平台默认 reasoning,short / medium / long prompt tiers,仅 Chat Completions | 中文 HTML · EN HTML · 中文 MD · EN MD · 报告目录 | 数据集 |
| 已发布 | openai/gpt-5.4-mini |
Infron vs OpenRouter | 4 x 50, streaming, routing sort: throughput / price / latency / ttft, platform-default reasoning, short / medium / long prompt tiers, Chat Completions only | 中文 HTML · EN HTML · 中文 MD · EN MD · 报告目录 | 数据集 |
| 已发布 | deepseek/deepseek-v4-pro |
Infron vs OpenRouter | 4 x 50,streaming,routing sort: throughput / price / latency / ttft,平台默认 reasoning,short / medium / long prompt tiers,仅 Chat Completions | 中文 HTML · EN HTML · 中文 MD · EN MD · 报告目录 | 数据集 |
| 已发布 | z-ai/glm-5.2 |
Infron vs OpenRouter | 4 x 50,streaming,routing sort: throughput / price / latency / ttft,平台默认 reasoning,short / medium / long prompt tiers | 中文 HTML · EN HTML · 中文 MD · EN MD · 报告目录 | 数据集 |
| 已发布 | deepseek/deepseek-v4-flash |
Infron vs OpenRouter | 4 x 50,streaming,routing sort: throughput / price / latency / ttft,平台默认 reasoning,short / medium / long prompt tiers,仅 Chat Completions | 中文 HTML · EN HTML · 中文 MD · EN MD · 报告目录 | 数据集 |
| 未发布 | llama/*, claude/*, gpt/*, other qwen/* |
Provider matrix expansion | Matched-payload A/B、streaming TTFT、provider attribution、cost breakdown | 待发布 | 待发布 |
首页按照 registry 的长期维度组织,后续实验数量增长时仍能保持清晰结构。
Infron 关于 LLM gateway 评估、路由与推理系统的研究论文,后续论文会继续追加到这里。
本项目把每一次 benchmark 都组织为可审计研究工件,而不是一次性的 dashboard 截图。
每一种 routing mode 都使用稳定 payload SHA256,每轮发送两次完全相同的 prompt。
只有 `sort/group/round` 下响应侧 prompt tokens 完全一致的 A/B 样本进入统计。
保留 usage、cost、TTFT、latency、provider 标识和 cache tokens 等响应字段。
报告、原始数据、代码快照和 checksum 一起提交。
Registry 已经为更大规模的模型与 provider 覆盖做好结构准备,同时保持同一证据标准。
继续加入 Qwen、Llama、Claude、GPT、Gemini、Mistral 以及垂直领域模型族。
追踪跨平台对比与上游 provider 行为,包括多轮重复实验中的 routing drift。
重复实验以观察 cache TTL、provider fallback、价格变化、尾延迟和吞吐稳定性。